arXiv q-fin
Sep 14, 2026
The Economics of Recursive Self-Improvement
The paper models recursive self-improvement as feedback loops whose net acceleration depends on the product of elasticities around each loop, while distinguishing narrow AI-R&D optimization from broader economically valuable capabilities. Its back-of-the-envelope calibration says current feedback is not yet strong enough for self-sustaining acceleration, although the estimated loops appear to be strengthening.
- In the model, net acceleration depends on the product of elasticities across each feedback loop.
- The paper's illustrative calibration suggests current feedback is below the threshold for self-sustaining acceleration.
Why it mattersThe framework turns a broad AI-capability narrative into parameters that labs could measure and disclose. The conclusion is model-dependent, uses limited public data, and should not be read as a forecast of when self-sustaining improvement will occur.
arXiv q-fin
Sep 11, 2026
Same Book, Different Fills: Partial Identification of FIFO Execution from Aggregate Order Books
Using matched Tokyo Stock Exchange L2 snapshots and L1 trades, the paper shows that passive-execution backtests can change materially under observationally equivalent FIFO cancellation rules. Across two instruments, front-versus-back cancellation changes preterminal completion by 7.39 to 8.01 percentage points and implementation shortfall by 0.384 to 1.010 basis points.
- Observationally equivalent aggregate-book paths produce economically different passive-execution outcomes under different cancellation allocations.
- The reported front-versus-back differences reach 7.39 to 8.01 percentage points for completion and 0.384 to 1.010 basis points for implementation shortfall.
Why it mattersBacktests built from aggregate books can conceal queue assumptions that materially change fill and cost estimates. The narrow two-instrument study supports sensitivity analysis rather than a universal correction and still needs broader replication.
arXiv q-fin
Sep 11, 2026
Diffusion models for dynamic volatility surface generation and data-driven hedging
The authors train sequential diffusion models on daily SPX option data from 2000 through 2023 to generate conditional volatility-surface paths and feed them into an optimization-based hedge. A post-trained variant adds static no-arbitrage penalties; the paper reports nearly eliminating such violations while reducing tail risk versus classical and data-driven baselines.
- The post-trained model is reported to reduce static no-arbitrage violations to nearly zero on the evaluated data.
- Diffusion-based hedges are reported to reduce tail risk and remain stable during the COVID-19 market disruption.
Why it mattersGenerative market scenarios are useful only if they respect financial structure and improve decisions. This paper connects scenario quality to hedging outcomes, but its backtests are author-reported, non-peer-reviewed, and do not establish live-trading performance.
arXiv cs.AI
Sep 11, 2026
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
Continual Search treats diagnosis of long-horizon agent failures as iterative evidence retrieval instead of a one-shot judgment. Across four existing benchmarks and a new 50-trial MegaRCA-Mix set, the authors report consistent gains; on MegaRCA-Mix, GPT-5.5 F1 rises from 0.349 to 0.498.
- The authors report that Continual Search improves root-cause attribution across multiple benchmarks and model families.
- On the 50-trial MegaRCA-Mix benchmark, reported GPT-5.5 F1 increases from 0.349 to 0.498.
Why it mattersOperational agent reliability depends on locating sparse causes in long execution traces. The result suggests search procedure can matter more than model tier, but the benchmark is small and the reported gains remain non-peer-reviewed.
arXiv q-fin
Sep 14, 2026
Resolution Is Not Settlement, Part II: Protocol Finality and Observed Redemption on Polymarket
Part II follows 108,638 exactly linked Polymarket conditions from preparation through protocol resolution and observed redemption. It finds that oracle finality, protocol finality, and holder realization are distinct; 99,283 conditions have an observed protocol-resolution event, while many cross-contract histories cannot be paired conservatively.
- The exact bridge links 108,638 conditions, of which 99,283 have an observed protocol-resolution event by the frozen snapshot.
- The formal results show that oracle finality does not identify protocol finality and protocol finality does not identify holder realization.
Why it mattersSeparating adjudication, payout recording, and actual redemption matters for settlement-risk measurement and prediction-market analytics. The study is descriptive and non-peer-reviewed, and redemption events alone do not identify the share of economic entitlement realized.
arXiv q-fin
Sep 14, 2026
Resolution Is Not Settlement, Part I: Oracle Adjudication and Semantic Governance on Polymarket
Part I reconstructs Polymarket oracle adjudication as an event sequence rather than one resolution timestamp. The frozen on-chain population contains 185,550 adapter-question instances and 504,332 decoded lifecycle events; exact metadata linkage covers 56.07% of questions, while unfinished histories remain right-censored.
- The study reconstructs 504,332 oracle lifecycle events across 185,550 adapter-question instances.
- Exact metadata linkage covers 56.07% of questions; the author does not substitute mechanism timestamps for unmeasured semantic-decidability timing.
Why it mattersPrediction-market research and risk systems can misstate timing and finality when they collapse proposals, disputes, resets, oracle finality, and adapter terminality into one field. The study is descriptive, single-author, non-peer-reviewed, and explicitly does not infer causal efficiency.
arXiv cs.AI
Sep 11, 2026
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
ZGCM-1 is a fully open 7B dense model trained for mathematical reasoning and agentic search with a 256K context window. The authors report roughly 4.2-times faster 16K pretraining time-to-loss, competitive results against much larger models on selected math and search suites, and release weights, checkpoints, training code, data recipes, and logs.
- The authors report a roughly 4.2-times improvement in 16K pretraining time-to-loss.
- The project releases weights across training stages, intermediate checkpoints, code, data recipes, and experiment logs.
Why it mattersA reproducible small-model training stack could make agentic-search research less dependent on closed frontier systems. The efficiency and benchmark claims are author-reported in a non-peer-reviewed preprint and need independent reproduction.
arXiv q-fin
Sep 13, 2026
Towards foundation models for insurance risk modelling
This review maps language, vision, geospatial, time-series, tabular, and scientific foundation models to insurance risk workflows and proposes evaluating predictive contribution, stability, and compliance before actuarial use. It highlights delayed outcomes, rare large losses, transfer failure, privacy, finer risk classification, and shared-provider dependence as core constraints.
- The paper proposes a process for connecting foundation-model outputs to actuarial calculations and testing prediction, stability, and compliance.
- It identifies delayed labels, rare losses, transfer risk, privacy, and shared-provider dependence as material evaluation problems.
Why it mattersInsurance foundation models could extract useful signals from claims and sensor data, but the paper's main contribution is a diligence framework rather than evidence of deployed performance. Its concentration and access-to-insurance warnings are directly relevant to model governance.